Skip to content

linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole - #582

Open
eKisNonos wants to merge 254 commits into
mainfrom
linux/go-guests
Open

eKisNonos wants to merge 254 commits into
mainfrom
linux/go-guests

Conversation

@eKisNonos

@eKisNonos eKisNonos commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Targets main. It also carries the 209 commits of #567 (linux/store-install), which come before it and hold the Linux personality these commits change, so it merges after #567 and shows those commits until #567 lands. What it adds is thirty-six commits.

Before this, a Linux-guest test image started no guest at all. Once it did, no guest could run a second thread: Go's threads called address zero, musl's pthread_create failed outright, and a thread's exit returned to it. A guest that did get threads reset the machine when it exited, left its threads running, and hung if one of them faulted. And nothing could wait: eventfd2 was not served, so every Go program with a timer died at its first one, while epoll_wait and a futex with a timeout returned at once or never.

What Linux guests can do now

  • Go programs start their runtime threads and run to a clean exit.
  • musl programs create pthreads, lock mutexes and join every thread.
  • A thread that faults ends the whole process, as on Linux; nothing is left running or waiting.
  • A fork from any thread gives the child that thread's registers and thread pointer.
  • Go's timers and poller work: time.Sleep, a timer woken early through Go's eventfd, and a pipe read through the poller to end of file.
  • musl programs wait as on Linux: timed condition variables, blocking eventfd reads, epoll_wait, poll and select timeouts, non-blocking pipes, several readers or a writer waiting on one pipe, edge-triggered epoll, and timerfd one-shot, periodic and absolute.
  • The descriptor ioctls, the scheduler calls, epoll_create, epoll_pwait2 and tgkill are served as Linux serves them to an unprivileged process on one CPU.

What changed

The boot guest starts

  • mk bf9d9dbf3: the Linux-guest test image (NONOS_LINUX_GUESTS=1) builds microkernel-desktop-gui with nonos-stark-attest, without first-boot setup. Under the setup profile every app, the personality included, waits for setup to exit, and setup waits for keys, so the unattended image never started its guest: 0 [APP-LINUX] or [LINUX] lines in 180 s. Every other build of nonos-mk-desktop-gui-prod is unchanged.

Threads start, run and end

  • foreign 3fc1ac19d: a clone child starts on a copy of the calling thread's registers, with rax 0, the new stack, and the parent's thread pointer unless a TLS is named. MkForeignThread takes the calling thread as a fifth argument; it must be supervised by the caller and in the same thread group as the target, so no guest receives another's registers. Zero keeps the fresh start. Go's child calls r12 and musl's calls r9, so both called address 0 before. A TLS is taken only when CLONE_SETTLS is set. abi/syscalls.toml records the new argument.
  • linux addec49af: clone writes the new tid where CLONE_PARENT_SETTID and CLONE_CHILD_SETTID ask. musl's thread-list lock stores its owner's tid, and a thread that never learned its tid locked the list as 0, which the lock reads as free.
  • linux 1e6b8a673: mprotect commits a PROT_NONE reservation it is asked to open. musl reserves every pthread stack with PROT_NONE and opens the usable part with mprotect; MkPeerProtect reprotects only pages that exist, so pthread_create failed. A backed region is reprotected as before; a reservation is backed with the asked protection, keeping any page the guest already touched, and recorded as backed so fork copies it. A span outside every region is refused with ENOMEM, as on Linux.
  • linux 10947aa96: a thread's exit ends the thread and never returns to it. The word it named through CLONE_CHILD_CLEARTID or set_tid_address is zeroed and one waiter woken, as Linux does. musl names its thread-list lock there and pthread_join waits for that lock, so every join hung before. Go's exitThread falls into INT3 when exit returns, which the fault report below would have turned into the end of the process.

A process ends whole

  • exit 5ab0a0912: a thread group's page tables stay until its last member leaves the process table. Release freed them when the leader was finalized, under a Go thread still running on them, and the machine reset: CR3 was the guest's, FS held the Go thread's TLS, CR2 was the IDT entry for vector 8. The tables now pass to a process still running on them, whose capability token is rebound to the ASID it now owns.
  • kill 874f3ecfc: MkKill admits the caller the foreign registry names as the target's supervisor. A thread's parent is its leader, not the supervisor, so every thread kill at a guest's exit returned EPERM and the threads ran on. A refused kill now prints [LINUX] kill refused: pid <n> outlives its process, errno <e>.
  • foreign 2311afc39: when a guest thread ends on a signal, and only when it ends itself so a supervisor's own kill cannot loop back, the kernel leaves a one-way notice for its supervisor, collected on its next MkForeignWait as a frame numbered FOREIGN_NR_DIED. The kernel reports; what a dead thread means is the personality's policy.
  • linux 959b57f98: the personality ends the process on that notice, as Linux ends a thread group on an unhandled fault, and reports the signal as 128+signo.

Fork

  • foreign ffc1b542b: a fork copies the calling thread's parked frame, and the kernel carries that thread's own thread pointer to the child. Before, every fork took the leader's registers and the personality's single fs_base, which is only the last thread to set one.

Waiting, as Linux waits

  • linux 69f3b48d4: a pipe has Linux's ends. The family notes, each time it lends its pipes to the guest being answered, which ends are still open anywhere in the family. From that note: a read end is readable while it holds bytes and hung up once no write end is left; a write end is writable while there is room and in error once no read end is left. An empty pipe with no write end reads as end of file, and a write no reader can take is refused with EPIPE. poll, epoll and select report hang-up and error unasked. Before, every pipe was readable and writable to all three.
  • linux 3e48edbaf: a descriptor keeps O_NONBLOCK. F_SETFL sets it, F_GETFL reports it with the access mode, pipe2 takes it, dup and fork carry it, and a non-blocking read of an empty pipe answers EAGAIN. Before, F_SETFL was dropped, which is how Go's runtime prepares every pipe for its poller.
  • linux bfb07892f: eventfd2 and eventfd are served. The counter is kept with the pipes as the family's, so dup, fork and every thread reach one. EFD_SEMAPHORE, EFD_NONBLOCK and EFD_CLOEXEC as Linux; all ones and a short buffer EINVAL. Go's runtime makes one at its first timer and threw when it could not.
  • linux 06a266ac5: epoll_wait and epoll_pwait wait for their timeout (an int of ms; negative waits until ready). A blocking eventfd read or write and a write to a full pipe wait too. A waiting call is left parked in its trap and tried again after every answer and at its deadline, so a write from another thread ends the wait at once. A wait on a socket or timer is looked at every 10 ms. A pipe write of up to PIPE_BUF goes in whole or waits. maxevents of zero or less is EINVAL.
  • linux 32d2a1a06: EPOLLET entries are reported as readiness rises and re-armed by an EAGAIN; EPOLLONESHOT reports once. epoll_ctl refuses as Linux: EBADF, EPERM for a regular file or directory, EINVAL, EEXIST, ENOENT. Go registers every pipe and socket with EPOLLET and EPOLLOUT, so a level-triggered report of an always-writable end would answer its poller at once, every time.
  • linux aa0087d75: a futex wait ends with ETIMEDOUT at its timeout. FUTEX_WAIT_BITSET (absolute, MONOTONIC or REALTIME), FUTEX_WAKE_BITSET, FUTEX_REQUEUE and FUTEX_CMP_REQUEUE are served. The value is compared as 32 bits, so musl's sign-extended -1 matches. musl's timed waits and condition variables, and Go's sysmon, use these.
  • linux b1ead15c5: two proof guests, gopoll and cwait (below).

More waiting, timers, descriptors and threads

  • linux b716a34af: any number of threads can wait to read one pipe. A read parked in one slot per process, so a second reader of an empty pipe took the slot and the first was never answered. Pipe reads now wait with the other parked calls; the family settles them after it reaps, so a writer that left with its process reads as end of file at once. A zero-byte read answers zero.
  • linux 439ed66c6: poll, ppoll, select and pselect6 wait for their timeout (poll's int of ms, ppoll's and pselect6's timespec, select's timeval; null or negative waits until ready; bad fields EINVAL). poll(NULL, 0, ms) sleeps. select writes its sets back only when something is ready, copies only the longs nfds covers, always empties the exception set, and empties all three when its time runs out. A negative poll descriptor is never ready.
  • linux 85b3d7724: a closed descriptor leaves every epoll interest list, as on Linux, so its number can be added again; before, that add was EEXIST and the old entry reported the new file under the old token. dup2 closes an open target first (a file's buffered bytes were lost, a socket's handle kept) and refuses a number past the table.
  • linux b9efc7f96: timerfd has Linux's semantics and is kept with the family like the eventfd counter, so dup and fork share it. A periodic timer keeps firing and a read returns how many times it fired; the clock is Linux's set (else EINVAL); TFD_NONBLOCK, TFD_CLOEXEC and TFD_TIMER_ABSTIME are honoured; the old setting is written when asked; timerfd_gettime is served; a blocking read waits for the firing; a wait watching a timer is woken when it fires.
  • linux 262c06280: FIONREAD (a pipe's bytes, a file's remainder), FIONBIO, FIOCLEX and FIONCLEX. Any other request is ENOTTY, as for a descriptor with no terminal or device behind it.
  • linux 8f778075b: a dup'd or inherited epoll descriptor carries the interest list it had; before, it started empty.
  • linux ff778c704: sched_getscheduler, sched_setscheduler, sched_getparam, sched_setparam, sched_get_priority_max and _min, sched_setaffinity, epoll_create, epoll_pwait2. Every thread is SCHED_OTHER at zero; SCHED_FIFO and SCHED_RR are EINVAL outside 1 to 99 and EPERM inside it, the range checked first as Linux checks it; a mask without the one CPU is EINVAL; pid arguments are mapped like kill's.
  • linux 9039b893d: tgkill's thread group and thread are mapped into the family's pids. gettid and getpid answer with the family's numbers, so tgkill(getpid(), gettid(), sig), which Go's runtime and glibc's pthread_kill use, was ESRCH.
  • linux 84e10f538, 7c3967f18, 468f5aa97, 2f0db3b25, a045bec06, 66b55faf8, afd96ef58, 24cafe103: the guest gopreempt, and seven more cwait parts, one per change above; cwait now runs every part and names each one that fails.

Cleanups

  • linux 289272374: delete call/clone.rs, a copy left behind when clone moved to call/spawn/; no module included it.
  • mm ce395f7ff: the [PF] demand fill line is printed only after a fill, not before the handler refuses the null page or the kernel half.

Evidence

q35 under TCG, one vCPU. Each guest is the boot guest of its own store.

Threads, faults and exit (the first thirteen commits)

guest proves line calls served unserved triple faults
gohello Go runtime threads, GC, clean exit [GO] hello PASS 271 0 0
goconc 8 goroutines on runtime threads [GO] conc PASS: 8 goroutines summed 3199960000 236 0 0
cthreads 8 musl pthreads, a mutex, 8 joins [C] cthreads PASS: 8 pthreads joined, summed 3199960000 87 0 0
threadfault a worker faults while main waits [LINUX] guest thread 51 ended on a signal; ending the process, then [LINUX] guest exited 7 0 0

threadfault's worker runs on a stack mapped read-write outright through clone(), so it does not depend on the pthread path. Its SURVIVED line, printed only if main outlives the faulting thread, never appears.

Waiting (the next seven commits)

Both guests pass on host Linux, which is what they are measured against. On one build carrying all twenty commits:

guest proves line calls served unserved triple faults
gopoll Go's timer and poller: a 50 ms sleep, a 10 ms timer set while the poller waits on a 3 s one, a pipe read through the poller [GO] poll PASS: 3 parts, 3454ms in all: slept 69 ms, the early timer ended in 18 ms, the pipe read to end of file in 339 ms 399 0 0
cwait musl waiting in 9 parts [C] cwait PASS: 9 parts in 951 ms: cond_timedwait ETIMEDOUT after 102 ms, a blocking eventfd read woken after 147 ms, epoll_wait timed out after 100 ms and woken after 153 ms, POLLHUP on a closed pipe, a full-pipe write waited 148 ms, edge-triggered counts as Linux 196 0 0

The same build ran the first four guests again: gohello 283 calls, goconc 253, cthreads 87, threadfault 7 with its death line; 0 unserved in each.

More waiting, timers, descriptors and threads (the next sixteen commits)

On one build carrying all thirty-six commits:

guest proves line calls served unserved triple faults
cwait 16 parts, the 9 above and 7 more [C] cwait PASS: 16 parts in 2313 ms; SCHED_FIFO at priority 1 errno 1 and a CPU-1-only mask errno 22; a 50 ms periodic timer fired 3 times in 180 ms, a one-shot read waited 103 ms, an absolute timer woke epoll after 105 ms 345 0 0
gopoll as above [GO] poll PASS: 3 parts, 3437ms in all 386 0 0
gopreempt a goroutine spinning with no call beside main, on one P no line in 360 s: this set does not reach a running thread, the gap named below 0

The same build ran the regression set: gohello 282 calls, goconc 248, cthreads 87, threadfault 7 with its death line; 0 unserved in each.

Each fix was also run with its change removed, on the same guest:

change removed guest result without it result with it
the supervisor clause in MkKill goconc 3 kill refused ... errno 1 lines; guest threads fault after guest exited 0 and 0
the kernel's death notice threadfault SURVIVED printed, no death line death line, no SURVIVED
the clear-tid write and wake cthreads no PASS, no FAIL, no exit in 241 s after the guest started PASS
the mprotect commit (the pthread build of threadfault, before the fix) threadfault [C] threadfault FAIL: no worker pthread_create succeeds (cthreads)
a parked wait tried again when something changes (only at its deadline instead) gopoll a 10ms timer set under a 3s one took 2823ms, FAIL 18 ms, PASS
the same cwait hangs in its blocking eventfd read; nothing after part 3 in 420 s PASS
eventfd2 served gopoll [LINUX] unserved nr=290, then fatal error: runtime: eventfd failed PASS
the futex timeout cwait hangs in cond_timedwait; no line in 420 s ETIMEDOUT after 102 ms
EPOLLET (reported level-triggered instead) cwait FAIL: edge-triggered counts (222, 1): the second look counts 2, not 1 PASS
the same gopoll PASS, with 1282 calls served: the poller answered at once while the pipe was open 399 calls
a pipe's hang-up cwait FAIL: end of file after the write end closed (0, 0): poll sees nothing POLLHUP
O_NONBLOCK kept by fcntl and pipe2 cwait hangs in the non-blocking pipe read; nothing after part 6 in 420 s EAGAIN at once
the personality before the next sixteen commits cwait eight parts pass, then it hangs in pipe-readers: the first of two blocked readers is never answered PASS
poll, ppoll and select waiting (answering at once instead), a closed descriptor leaving epoll, periodic timers, FIONREAD, sched_getscheduler, the tgkill mapping, all six removed on one build cwait FAIL: 6 parts failed, 10 passed: poll and ppoll returned after 2 ms, the re-add was EEXIST, the periodic timer counted 1, FIONREAD failed, [LINUX] unserved nr=145, tgkill ESRCH 16 parts pass
the epoll list carried through fork cwait FAIL: epoll list through fork (13, 256): the child finds nothing ready and exits 1 PASS

Each of these ran on its own build with only that change removed, except the six that fail different cwait parts, which were removed together. gopoll passes with the hang-up removed, as its writer closes before its reader waits, and with O_NONBLOCK dropped, as Go's pipe reads then park instead of polling.

The store settled in 12.8 to 84.2 s across the 22 guest boots for the first thirteen commits, in 38.0 to 104.8 s across the 16 for the next seven, and in 17.2 to 117.7 s across the 10 for the sixteen after; the personality waits up to 300 s.

Checks

  • x86_64 kernel crate: 0 errors; 43 warnings, the same set before and after.
  • check_stubs, check_allows, check_dark_features: 0 new sites. check_syscall_abi: 111 published syscalls reach a handler.
  • check_unreachable: 1 new site, has_children, which linux/store-install reports too; this branch has 1257 sites against its 1258.
  • Each of the first 13 commits builds on its own: the kernel (microkernel-core) and the personality were checked at every one, with 0 errors. The next twenty-three change only the personality and the guests; the personality was checked at each with 0 errors and 0 warnings, and each guest passes on host Linux.

Not done here

  • A PROT_NONE reservation is not enforced: the kernel demand-fills any guest page on first touch, so a guard page does not stop a stack overflow.
  • mprotect on a backed region does not update the recorded protection, so a fork after it gives the child the protection the region was mapped with.
  • A fork does not copy a reserved region the guest has touched, since only backed regions are copied.
  • The leader's own exit (not exit_group) still ends the whole process; on Linux the others would run on.
  • A caught signal reaches a guest thread only when it next makes a call. A thread parked in a wait sees it when the wait ends, and a thread running user code never does, so Go cannot preempt a goroutine that spins without a call (gopreempt, above). The kernel change that stops a running thread for its supervisor follows this set.
  • A caught signal sent with kill to a child process is queued in the sender, not the child, so the child's handler never runs; an uncaught one ends the child as it should.
  • A write to a pipe with no reader is EPIPE, but SIGPIPE is not raised.
  • A wait on a socket is looked at again every 10 ms; nothing tells the family sooner that a socket became ready.
  • A dup'd or inherited epoll descriptor gets a copy of the interest list; on Linux both share one, so a change made on one side after the dup or fork is not seen on the other.
  • TFD_TIMER_CANCEL_ON_SET is accepted and has no effect.

senseix21 and others added 30 commits September 24, 2026 22:00
IconId::Store points table.rs at assets/icons/store.a8, which only
existed on the app-store branch, so every build of this branch failed
to read it. Add the mask and its SVG source here so the table stands
on its own.
text::line returns the drawn width, so the bare match evaluated to i32
where the function body expects (), failing the capsule build.
nonos-data/marketplace/index.bin has no make rule, so naming it as a
hard prerequisite failed every build on a checkout without it (CI:
No rule to make target). Wrapping it in $(wildcard) keeps the rebuild
on a newer catalogue where it exists and drops the prerequisite where
it does not.
The branch's manifest predated the switch to in-process Ed25519 and
dropped the dependency while verify/crypto.rs imports it, so the
capsule failed with an unresolved import. Restore main's manifest and
add only the app_skeleton dependency boot_index.rs needs.
Twenty-five syscalls were published as caps = ["valid_token"] while the
cap table demands a hardware, dev-root or time capability. Twenty-two take
any one of Admin or a hardware capability, now published with caps_any; the
dev-root and time calls need one capability, published with caps.
The cap table gates MDRO with MDRQ and MDRC on can_enrol_dev_root. The old
syscall caps check cannot resolve that predicate and wants valid_token, so
this fails it until abi/caps-check-fail-closed lands; the fixed check passes.
Every lane was pinned to -accel hvf -cpu host, which only macOS has, so
no Linux host could boot an image. KVM when /dev/kvm opens read-write,
hvf on macOS, TCG otherwise, the rule the boot matrix already uses; the
display and audio backends follow the host too.
The trailer's magic picked the verifier for every root, so a Pedersen
trailer was checked against the vendor root too. A local root's leaf is a
commitment to a secret this kernel holds; the vendor root's is not, so there
the trailer may no longer choose the weaker proof.
MkLocalSign let a LocalSign holder prove any capability it held, including
LocalSign itself, so one signer could hand out the right to sign. A local
proof now names nothing beyond AMBIENT_CAPS. crypto_proofs checks a minted
trailer and its refusals against the kernel's own verifier files.
An Alpine package index is signed that way, and the capsule offered SHA-1
only without the prefix, so the index could not be checked at all. The
rsa crate rebuilds the whole padded block and compares it. Nothing here
signs, so offering SHA-1 verification mints nothing new with it.
resolve.rs names super::root, which the crate never mounted, so it did not
build and nothing noticed because no workflow ran it. It mounts root.rs,
follows the rename of absolute to visible, adds the /linux confinement and
Phdr::file_range tests, and joins the proof-crate matrix.
A signature over a package covers the compressed bytes of one member, so a
verifier needs to know where each one starts and ends. members() refuses a
file with any byte that belongs to no member, where gunzip ends the stream
and ignores the rest.
The index's .SIGN.RSA entry is checked against Alpine's x86_64 keys by the
crypto service, the package's control member against the index's C: SHA-1,
and its data against the control member's datahash. Only a Verified value
reaches the store, so unauthenticated bytes are refused, not kept unvouched.
Without the operator-key rotation: NOX_OPERATOR_V1 keeps main's key, now
read from .keys/marketplace_operator_ed25519.pub. The capsule embedded an
index nothing built, so tools/nonos-market-index writes one, signed and
verified when the operator seed is present and empty otherwise. nonos-mk
moves to 78dae45 for the zk_trailer_hash field; the icon table is 49 long.
A release naming x86_64-linux counted as having its attestation, so any
release could claim the exemption for itself. It now needs the linux.
namespace too, which is where the store routes it: to the installer that
authenticates the bytes before the machine mints their proof.
Conflicts were both sides adding: the install and app-store capsules sit
together in Cargo.toml, mk and userspace, and init/mod.rs names the
install queue once. The init loop and install queue are taken as the PR
wrote them; the commit after this replaces its wake path.
The scheduler takes every ready process's priority lock from the timer
interrupt. wake.rs and boost_init_for_drain took init's with interrupts on
from syscall context, so a tick inside either spun forever on one CPU. The
install queue now raises through the guarded setter the window queue uses.
MkAppInstall took a package name and the store's own readiness flag. It now
takes a listing and release; init asks the market for readiness and the
release's package hash, and the installer refuses bytes of any other BLAKE3.
The index keeps each record's D: and p: lines, so dependencies come too.
…each

CryptoMachineKey derives the machine key for any label a Crypto holder
names, so a key the kernel keeps for itself needs a label no syscall can
ask for. Kernel labels start with a zero byte, and the syscall refuses any
label that does.
The local signing identity was random each boot, so consent was too. It is
now derived from the machine key, and first-boot setup, which alone holds
EnrolDevRoot, grants the local root as a named step and keeps a token only
this machine can make; later boots restore it. The desktop profile now
includes setup and the market, and builds every capsule it embeds.
MkAppLaunch queues a run for init, which spawns the personality to start
the program the installer recorded outside /linux. Whether it may start is
the exec gate's answer. The store drops its console-code enrolment, which
it never held the capability for, and gains an o key to open.
Every boot resolved validator.nymtech.net through net.dns, so the first
query of a session went out in the clear naming the service in use. The
bootstrap set is pinned by address in this capsule's attested image; the
name is kept for the certificate check and never resolved.
The installer connected to dl-cdn.alpinelinux.org, which the socket service
resolved in the clear. It now takes a mirror by address, Alpine's CDN by
default or NONOS_ALPINE_MIRROR at build, sends Alpine's name as the Host
line, and refuses a host that is not a literal. The signatures still decide.
The generator hashed packages from v3.21 while the installer downloads
from v3.20, so every Linux listing pinned bytes the installer never sees
and the kernel's package-hash check would refuse every install.
The spawn installs an empty token, and nothing on the console said so.
The line reads the bits back from the process table rather than printing
the value it meant to install, so it is evidence and not a restatement.
The four answered at once whatever their timeout, so a program waiting on
descriptors spun, and poll(NULL, 0, ms), the idiom for a short sleep, did
not sleep. select also narrowed the caller's sets even when nothing was
ready, and copied a whole 1024-bit set whatever nfds was. poll counted an
entry with a negative descriptor, which Linux ignores and programs use to
switch an entry off, as closed and so ready.

Each now waits with the other parked calls until something it watches is
ready or its timeout passes: poll's int of milliseconds, negative for none;
ppoll's and pselect6's timespec and select's timeval, null for none, with
negative or out-of-range fields refused with EINVAL. A timeout of zero
only looks. select writes its sets back only when something is ready,
copies only the longs nfds covers, always empties the exception set, and
empties all three when its time runs out, as Linux does. A negative poll
descriptor is never ready. A wait watching a socket or a timer is looked
at again every 10 ms, now for poll and select as for epoll.
Two parts added to cwait. Two threads block reading one empty pipe and both
must be answered once bytes arrive. poll, ppoll and select each wait out a
100 ms timeout with nothing ready, poll(NULL, 0, 50) sleeps, and a select
with no timeout is woken by another thread's write, with its set narrowed
to the ready descriptor. Both pass on host Linux.
…aces

A closed descriptor stayed in every epoll interest list. Linux drops it, so
a program that closes without EPOLL_CTL_DEL and gets the same number back
from its next open adds it again; here that add was refused with EEXIST,
and the old entry reported the new file under the old token. dup2 onto an
open descriptor overwrote it without closing it, so a file's buffered bytes
were lost and a socket's handle kept, and a target number past the table
grew the table to reach it.

close now drops the descriptor from the guest's interest lists. dup2 closes
an open target first, as Linux does, and refuses a number past the table
with EBADF.
A twelfth part: a watched pipe end is closed without EPOLL_CTL_DEL, a new
pipe takes its number, and adding it again succeeds with no stale event.
dup2 to a number past the table is EBADF. Passes on host Linux.
A timerfd fired once and forgot its interval, so a periodic timer stopped
after its first read, and a read always answered one. timerfd_create
ignored its clock and its TFD_NONBLOCK and TFD_CLOEXEC flags,
timerfd_settime ignored TFD_TIMER_ABSTIME and never wrote the old setting,
timerfd_gettime was not served, and a read of a timer that had not fired
answered EAGAIN even to a blocking descriptor. The expiry lived in the
descriptor, so dup and fork gave a timer that never fired.

A timer is now an object the family keeps, lent with the pipes and eventfd
counters, so every descriptor onto it sees one timer. It is made on a
clock Linux names (anything else is EINVAL) with its flags; set relative or
absolute on its own clock, one-shot or periodic, with the old setting
written when asked; read for how many times it has fired since the last
read, stepping a periodic timer past them; read blocking until it fires
unless the descriptor is non-blocking; and reported readable once it has
fired. A wait watching a timer is woken when it fires rather than on a
10 ms look.
A thirteenth part: an unarmed non-blocking timer reads EAGAIN, a 50 ms
periodic timer counts its firings over 180 ms and reports its interval, a
one-shot read blocks until it fires, and an absolute time 100 ms ahead
wakes an epoll wait. Passes on host Linux.
gopreempt spins a goroutine with no call in it beside main, on one P. Go
moves such a goroutine off the CPU only by sending SIGURG to its thread
while it runs, so main waking from a 20 ms sleep and a garbage collection,
which stops the world, both depend on that signal landing. On host Linux
main runs again after 25 ms and the collection ends by 45 ms, with the
tgkill deliveries visible under strace.
ioctl answered ENOTTY to every request, so a program asking how many bytes
a pipe or file held, setting a descriptor non-blocking the BSD way, or
marking it close-on-exec without fcntl was refused.

FIONREAD answers what a pipe's read end holds, or what is left of a file
past its offset; FIONBIO sets or clears O_NONBLOCK from the int it is
given; FIOCLEX and FIONCLEX set and clear close-on-exec. A socket's
FIONREAD, and every device request, is still ENOTTY: there is no terminal
or device behind a descriptor here, which is also how isatty says no.
A duplicated or inherited epoll descriptor started with an empty interest
list, so a child that waited on the epoll its parent set up before fork
waited on nothing. The list is now copied with the descriptor. Linux shares
one list between them; the copy holds what was registered at the dup or
fork, and a change made afterwards on one side is not seen on the other.
A fourteenth part: FIONREAD counts three bytes in a pipe, FIONBIO makes its
read end answer EAGAIN once drained, FIOCLEX sets close-on-exec, and a
forked child finds a ready entry in the epoll list its parent built.
Passes on host Linux.
sched_getscheduler, sched_setscheduler, sched_getparam, sched_setparam,
sched_get_priority_max and _min and sched_setaffinity were not served, nor
epoll_create and epoll_pwait2, so a program that sets a thread's policy or
pins it, or a libc that reaches for the older or newer epoll form, died
there.

They are answered as Linux answers an unprivileged process on the one CPU
the guest is shown: every thread is SCHED_OTHER at priority zero; asking
for SCHED_OTHER, BATCH or IDLE at zero is accepted and changes nothing,
since scheduling is the kernel's; asking for FIFO or RR is EINVAL outside
the priority range 1 to 99 and EPERM inside it, the range checked first as
Linux checks it, since no guest holds the privilege Linux asks for; the
priority ranges are Linux's. An affinity mask that includes the one CPU is
accepted and one that leaves it out is EINVAL. A pid argument is mapped
from the guest's namespace like kill's, and one outside the family is
ESRCH. epoll_create checks its size is positive and makes a list;
epoll_pwait2 waits like epoll_pwait with a timespec timeout.
A fifteenth part: the policy is SCHED_OTHER, SCHED_FIFO at priority zero is
EINVAL, SCHED_OTHER is accepted, the real-time range is 1 to 99, CPU 0 can
be pinned, epoll_create refuses a size of zero, and epoll_pwait2 waits its
50 ms timespec. SCHED_FIFO at priority one and a CPU-1-only mask depend on
privilege and CPU count, so they are printed, not checked. Passes on host
Linux.
gettid and getpid answer with the family's own numbers, and kill and tkill
map theirs back, but tgkill's were passed through as they came. A thread
signalling itself with tgkill(getpid(), gettid(), sig), which is how Go's
runtime preempts a goroutine and how glibc's pthread_kill and raise reach a
thread, named numbers no kernel thread has and was answered ESRCH.

Both of tgkill's pid arguments are now mapped, and one outside the family
is ESRCH, as for kill.
A sixteenth part: a caught SIGUSR1 sent with tgkill(getpid(), gettid())
runs its handler, and a thread the family does not have is ESRCH. Passes
on host Linux.
cwait stopped at its first failing part, so a run with more than one
change removed showed only the first. Every part now runs, each failure is
printed, and the last line counts them. A hang still stops it at the part
it hangs in.
A signal reached a guest thread only as the answer to a call it had made,
so a thread running its own code never received one. Go preempts a
goroutine that spins without a call by sending SIGURG to its thread, and a
guest is shown one CPU, so such a goroutine held the only P for good: no
other goroutine ran and a garbage collection never stopped the world.
SIGALRM, SIGINT and any other signal for a busy thread waited the same way.

MkForeignInterrupt lets a supervisor mark one of its guest threads. The
timer trampoline, after a tick that interrupted user mode, parks a marked
thread with its whole register file as a frame numbered NR_INTERRUPTED and
wakes the supervisor; the thread sleeps as a parked call does. A signal
answer rewrites the trampoline's frame to enter the handler, keeping the
thread's FPU state for its return as a delivered call does; any other
answer lets the thread run on exactly where it was. A thread already
parked in a call is not marked, since that call's answer can carry the
signal. Only the supervisor recorded for the thread may mark it; the mark
is dropped with the thread; the handler's context is checked as for any
signal answer. The trampoline's frame is its 160 bytes and no more: it is
read and written as those 20 words, never as the whole SavedUser, whose
TLS words would lie past the top of the kernel stack. A tick with nothing
marked costs one load. The libc gains mk_foreign_interrupt and
FOREIGN_NR_INTERRUPTED.
A caught signal raised for another thread waited in the queue until that
thread next made a call, so a thread running its own code never saw it.

Raising one now also asks the kernel to stop the target at its next tick.
The stopped thread arrives as a FOREIGN_NR_INTERRUPTED frame and is
delivered what is pending, with its handler returning to the thread's own
rax since no call is being answered; with nothing pending it runs on where
it was. A thread parked in a call, the caller included, still gets the
signal with that call's answer.
cpreempt spins a thread on a flag only its SIGUSR1 handler sets, sends it
SIGUSR1 with pthread_kill, and joins it. The thread makes no call while it
spins, so the join returns only if the handler runs inside the spin. It
passes on host Linux, pinned to one CPU as well.
@eKisNonos eKisNonos changed the title linux: run Go and musl threads in a guest, and end the guest whole linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole Sep 28, 2026
The frame a handler is entered on wrote the 18 saved registers at
ucontext + 48 and the blocked mask right after them. Linux, musl, glibc
and Go all read uc_mcontext at ucontext + 40 and uc_sigmask at + 296
(measured with offsetof on musl and glibc). Every register a handler
read from its context was therefore the one before it: Go's preemption
handler reads rip and rewrites rip and rsp, so it would have read the
saved rsp as the pc and resumed into garbage. The C handlers so far
never read their context, and the frame tests only read it back the way
it was written, so nothing showed it.

The registers now start at + 40 and the mask sits at + 296. rt_sigreturn
reads the same constant, so a frame still round-trips, and a new test
pins r8, rsp, rip and the mask to the offsets musl and glibc use. With
the old offset that test fails and the four round-trip tests pass.
… handlers on it

sigaltstack read back an empty stack and ignored the one a program set,
so every handler was entered on whatever stack the thread was using.
Go installs every handler with SA_ONSTACK and gives each thread its own
signal stack; its handler checks that it runs on that stack or on g0's,
and a frame on a goroutine's stack sends it into a path that waits for a
spare M a program without cgo never has. That is where gopreempt stopped:
one SIGURG delivered to tid 50, then no further line in 420 s.

Each thread's alternate stack is now kept, and sigaltstack answers as
Linux's do_sigaltstack: the old setting is what was there before the
call, SS_ONSTACK while the thread runs on it and SS_DISABLE when none is
set; changing it from on it is EPERM, a stack under MINSIGSTKSZ (2048) is
ENOMEM, and a mode other than 0, SS_ONSTACK or SS_DISABLE is EINVAL.
SS_AUTODISARM is named unserved and answered EINVAL, as a Linux before 4.7
answers it. A handler with SA_ONSTACK is entered at the top of the
alternate stack when the thread is not already on it, and below the
interrupted rsp when it is; a frame that would run off the bottom is not
written. uc_stack carries the stack and its flags for the interrupted rsp.
A fork gives the child the forking thread's stack, a thread's exit drops
its own, and execve clears them all.
A seventeenth part sets a 16 KiB alternate stack through the raw call,
so the answers are the kernel's and not musl's own checks, and raises a
SIGUSR2 caught with SA_ONSTACK. It passes only if: none set reads back
SS_DISABLE; 1024 bytes is ENOMEM and flags 4 is EINVAL; the stack set
reads back with flags 0; the handler's own local is on that stack while
the saved rsp is not; inside the handler the stack reads SS_ONSTACK and
changing it is EPERM; uc_stack names it with flags 0; and SS_DISABLE
turns it off again. On the container's Linux it prints
"cwait altstack ok: ... flags inside 1" and "cwait PASS: 17 parts".
MkForeignInterrupt marks a running thread so that the next tick in user
mode stops it and hands it to its supervisor. When the thread makes a
call before that tick, the call's answer is where the supervisor
delivers the signal, but the mark stayed, and a later tick stopped the
thread once more with nothing to deliver. The instrumented gopreempt run
showed it: after "signal 23 to tid 50" on a call's return, the next tick
took tid 50 again and the supervisor answered it with a plain resume.

A thread's mark is now dropped as it traps into a call. The check costs
one atomic load while nothing is marked.
clone returned the new thread's tid in the guest's pid namespace, but
wrote the kernel's pid for it to the CLONE_PARENT_SETTID and
CLONE_CHILD_SETTID words. musl keeps the parent's copy as the thread's
own tid and hands it to tkill, so pthread_kill asked for a tid the guest
does not have and got ESRCH. An instrumented cpreempt showed it,
"[C] dbg kill rc 3", then a join that waited for a signal never sent.
gettid and Go's tgkill were right, since both use the translated number.

The family now writes those words once the reply has been put in the
guest's terms, with the same number clone returns. clone itself keeps
recording CLONE_CHILD_CLEARTID, which is keyed by the kernel's pid.
An eighteenth part starts a thread that yields until a SIGUSR1 handler
sets its flag, then sends pthread_kill to it and joins. It passes only if
pthread_kill answers 0 and the handler ran. The thread gives up after
3 s, and sooner when the kill failed, so a lost signal fails the part
and does not hang it. On the container's Linux it prints
"cwait pthread-kill ok ... handled 1" and "cwait PASS: 18 parts".
eKisNonos added a commit that referenced this pull request Sep 30, 2026
#582 gained nine commits on 30 September. Four are already here as the
same commits. The other five are earlier forms of 07d9c92, d7b2019,
f5fe356, 2b5c726 and 5b40bbc, which this branch carries in the split
form linux/next-waits gave them. The tree here already holds all of it, so
the merge keeps it as it is, and #582 and this branch now merge into main
in either order.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants